Papers with toxicity model
Perturbation Sensitivity Analysis to Detect Unintended Model Biases (D19-1)
Copied to clipboard
| Challenge: | Recent research shows that data-driven NLP models may inadvertently capture, reflect and sometimes amplify various social biases present in the language data they are trained on. |
| Approach: | They propose a generic evaluation framework that detects unintended model biases related to named entities and requires no new annotations or corpora. |
| Outcome: | The proposed framework detects unintended model biases related to named entities and requires no new annotations or corpora. |